Papers with data construction process

4 papers
Explainable CED: A Dataset for Explainable Critical Error Detection in Machine Translation (2024.naacl-srw)

Copied to clipboard

Challenge: Existing studies of critical error detection lack content addressing the causes of catastrophic errors.
Approach: They propose a dataset that introduces the attributes of error explanation and correction regarding critical errors.
Outcome: The proposed dataset reduces time costs and mitigates human annotation bias.
MedDistant19: Towards an Accurate Benchmark for Broad-Coverage Biomedical Relation Extraction (2022.coling-1)

Copied to clipboard

Challenge: Relation extraction in the biomedical domain is challenging due to the lack of labeled data and high annotation costs.
Approach: They propose to use distant supervision to pair knowledge graph relationships with raw texts to tackle the scarcity of annotated data and to validate their results.
Outcome: The proposed benchmarks are more accurate and consistent with existing benchmarks and show that there is no train-test leakage.
Instance Regularization for Discriminative Language Model Pre-training (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies have optimized independent strategies of ennoising or denosing . Existing methods treat training instances equally throughout the training process .
Approach: They propose to use ennoising and denoising to train discriminative pre-trained language models . they propose to model the complexity of restoring the original sentences from corrupted ones .
Outcome: Experimental results show that the proposed method improves pre-training efficiency, effectiveness, and robustness.
Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: In real-world scenarios, user instructions often contain soft constraints, which are semantically related and cannot be rule-based verified, posing challenges for large language models.
Approach: They propose a pipeline to construct datasets with high-quality outputs for instructions containing soft constraints automatically and use Direct Preference Optimization (DPO) as the training method.
Outcome: The proposed model improves the LLMs' soft constraint following ability by using direct preference optimization (DPO) and constraint quantity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations